FIX: Treating simulated conversations properly - #2629
Open
Richard Lundeen (richlundeen) wants to merge 3 commits into
Open
Richard Lundeen (richlundeen) wants to merge 3 commits into
Richard Lundeen (richlundeen) wants to merge 3 commits into
Conversation
Prevent nested simulated-conversation attacks from creating standalone attack history rows while retaining their conversation lineage on the primary attack. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Roman Lutz (romanlutz)
requested changes
Sep 11, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Roman Lutz (romanlutz)
approved these changes
Sep 12, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Simulated conversations are generated by a nested
RedTeamingAttackbefore the primary attack runs. That nested attack was persisted as an independent attack result, so the GUI could show internal preparation work as a separate history row. For example, a scenario against GPT-5.1 could display an extraRedTeamingAttackthat targeted the Grok model used to generate the simulated prefix.The fix prevents the nested preparation attack from creating its own attack-result row while keeping its messages in memory.
AttackResultnow tracks simulated conversations as typed related conversations, similar to how it tracks pruned conversations. It links the simulated target conversation asPREPARATIONand retains the generator conversation asADVERSARIAL; these diagnostic conversations remain tracked but are not selectable as normal chat conversations.